Natural Language Processing and Biological Methods
نویسندگان
چکیده
During the 20th century, biology—especially molecular biology—has become a pilot science, so that many disciplines have formulated their theories under models taken from biology. Computer science has become almost a bio-inspired field thanks to the great development of natural computing and DNA computing. From linguistics, interactions with biology have not been frequent during the 20th century. Nevertheless, because of the “linguistic” consideration of the genetic code, molecular biology has taken several models from formal language theory in order to explain the structure and working of DNA. Such attempts have been focused in the design of grammar-based approaches to define a combinatorics in protein and DNA sequences (Searls, 1993). Also linguistics of natural language has made some contributions in this field by means of Collado (1989), who applied generativist approaches to the analysis of the genetic code. On the other hand, and only from theoretical interest a strictly, several attempts of establishing structural parallelisms between DNA sequences and verbal language have been performed (Jakobson, 1973, Marcus, 1998, Ji, 2002). However, there is a lack of theory on the attempt of explaining the structure of human language from the results of the semiosis of the genetic code. And this is probably the only arrow that remains incomplete in order to close the path between computer science, molecular biology, biosemiotics and linguistics. Natural Language Processing (NLP) –a subfield of Artificial Intelligence that concerns the automated generation and understanding of natural languages— can take great advantage of the structural and “semantic” similarities between those codes. Specifically, taking the systemic code units and methods of combination of the genetic code, the methods of such entity can be translated to the study of natural language. Therefore, NLP could become another “bio-inspired” science, by means of theoretical computer science, that provides the theoretical tools and formalizations which are necessary for approaching such exchange of methodology. In this way, we obtain a theoretical framework where biology, NLP and computer science exchange methods and interact, thanks to the semiotic parallelism between the genetic code and natural language.
منابع مشابه
روش جدید متنکاوی برای استخراج اطلاعات زمینه کاربر بهمنظور بهبود رتبهبندی نتایج موتور جستجو
Today, the importance of text processing and its usages is well known among researchers and students. The amount of textual, documental materials increase day by day. So we need useful ways to save them and retrieve information from these materials. For example, search engines such as Google, Yahoo, Bing and etc. need to read so many web documents and retrieve the most similar ones to the user ...
متن کاملProcessing and stabilization of Aloe Vera leaf gel by adding chemical and natural preservatives
Background and objectives: Aloe vera has been used as a medicinal herb for thousands of years. Aloe vera leaves can be separated into latex and gel which have biological effects. Aloe gel is a potent source of polysaccharides. When the gel is exposed to air, it quickly decomposes and decays and loses most of its biological activity. There are various processin...
متن کاملCorpus based coreference resolution for Farsi text
"Coreference resolution" or "finding all expressions that refer to the same entity" in a text, is one of the important requirements in natural language processing. Two words are coreference when both refer to a single entity in the text or the real world. So the main task of coreference resolution systems is to identify terms that refer to a unique entity. A coreference resolution tool could be...
متن کاملTowards Incorporating Scientific Literature into Biological Algorithms
The Tutorial This tutorial is a practical introduction on applying natural language processing (NLP) to biological research. It focuses on basic methodology used in current research efforts both in the published literature and in our lab. This tutorial exposes attendees to a broad range of work in the field, and at the end they will understand the technologies applied to solve NLP problems in b...
متن کاملThe Effect of Processing of Corn Silage with Schizophyllum Commune on Chemical Composition, Ruminal Degradability and in Vitro Gas Production
Extended Abstract Introduction and Objective: To overcome the problems caused by animal feed shortages, efforts are made to increase the availability of nutrients and their digestibility, such as improving the nutritional value of forage plants through biological processing. Schizophyllum commune is an edible fungus of the basidiomycete’s family that has been used for biological processing of ...
متن کاملUsing Generalized Language Model for Question Matching
Question and answering service is one of the popular services in the World Wide Web. The main goal of these services is to finding the best answer for user's input question as quick as possible. In order to achieve this aim, most of these use new techniques foe question matching. . We have a lot of question and answering services in Persian web, so it seems that developing a question matching m...
متن کامل